- Title
- Principal components analysis in stylometry
- Creator
- Craig, Hugh
- Relation
- ARC.DP160101527 http://purl.org/au-research/grants/arc/DP160101527
- Relation
- Digital Scholarship in the Humanities Vol. 39, Issue 1, p. 97-108
- Publisher Link
- http://dx.doi.org/10.1093/llc/fqad083
- Publisher
- Oxford University Press
- Resource Type
- journal article
- Date
- 2024
- Description
- Principal components analysis (PCA) has been one of the staple methods used in stylometry. In a 2021 article, Pervez Rizvi casts doubt on this method and argues that some widely cited results based on it should be set aside. In the current article, I show that none of Rizvi's theoretical claims or experimental results stand up to examination. Rizvi argues that discarding the principal components beyond the first two makes the method unreliable, but permutation testing of PCAs shows that the top components in these trials are significant and robust, and the results across many experiments show the combination of the first and second component to be effective in classification. Rizvi argues that PCA components must be treated separately, and much of his critique of the PCA method is based on this standpoint, but this is not the practice in the work presented in the publications he cites or in the wider literature. Rizvi is unable to replicate a chart in an article by Craig, but his replication, unlike the original, does not account for the widely varying sizes of samples in his data. The current article shows that Rizvi's claims are misguided and that using PCA in the Burrows tradition to find and formalize authorial discriminations in text samples from plays of the Shakespearean era is efficacious and robust.
- Subject
- stylometry; principal components analysis (PCA); Shakespeare; data analysis
- Identifier
- http://hdl.handle.net/1959.13/1505793
- Identifier
- uon:55742
- Identifier
- ISSN:2055-7671
- Language
- eng
- Reviewed
- Hits: 795
- Visitors: 795
- Downloads: 0